Skip to content

nvshmem4py - #674

Draft
jeffhammond wants to merge 12 commits into
ParRes:mainfrom
jeffhammond:nvshmem4py
Draft

jeffhammond wants to merge 12 commits into
ParRes:mainfrom
jeffhammond:nvshmem4py

Conversation

@jeffhammond

Copy link
Copy Markdown
Member

WIP

jeffhammond and others added 12 commits August 26, 2025 14:06
…d-to-end

cuda-python/cuda-core promoted its Device/system API out of the
cuda.core.experimental namespace into stable cuda.core sometime after
these scripts were written, and renamed system.num_devices (an
attribute) to system.get_num_devices() (a method) along the way. Both
PYTHON/hello-nvshmem.py and PYTHON/nstream-cupy-nvshmem.py were still
importing from the old experimental namespace, so neither could even
import against a current install.

Set up a local .venv (PRK/.venv, gitignored) with nvshmem4py-cu13 +
cuda-python + cupy-cuda12x + mpi4py -- nvshmem4py itself isn't on the
system Python and isn't packaged in this repo's own dependency tooling,
so a project-local venv keeps it isolated. Verified both scripts run
correctly under `mpirun -n 2` with this venv: hello-nvshmem.py completes
and prints OK on both PEs, and nstream-cupy-nvshmem.py runs the full
STREAM triad and validates.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The prior commit's "hack to disable jit?" (409c868) added a
nvshmem.barrier() after every loop iteration because timings looked
wrong without it. Investigated with a real, working nvshmem4py install
(see previous commit): the actual bug is that CuPy operations issued
with no explicit stream run on CuPy's own default stream, a different
stream than the cuda.core `stream` object used for NVSHMEM's
barrier/sync -- so stream.sync() was never actually waiting for the
triad kernels to finish, and the CPU-time timer (see below) raced far
ahead of the real GPU work. The per-iteration barrier "fixed" the
symptom only by forcing the host to block/spin on every iteration, which
happened to make time.process_time() advance too -- at the cost of also
timing NVSHMEM barrier overhead every iteration instead of just the
compute.

Also switched the timer from time.process_time() (CPU time) to
timeit.default_timer() (wall clock): process_time is the right choice
for the CPU-bound sibling scripts (nstream.py, nstream-numpy.py, etc.,
where process/wall time track closely), but is fundamentally wrong for
this GPU-async one -- it barely advances while the host blocks on the
device.

Fixed by binding CuPy's current stream to the same cuda.core stream
NVSHMEM uses (cupy.cuda.Stream.from_external(stream), which cuda.core's
Stream supports natively via the __cuda_stream__ protocol -- no
deprecated ExternalStream needed), removing the per-iteration barrier,
and switching to a wall-clock timer.

Verified: `mpirun -n 2 python3 PYTHON/nstream-cupy-nvshmem.py 10
16777216` and a larger `20 134217728` run both validate correctly and
report bandwidth in the range of real H100 HBM bandwidth (944 GB/s and
2.7 TB/s respectively, scaling up sensibly as problem size grows past
launch-overhead-dominated territory) -- versus ~16 GB/s before this fix,
which was dominated by per-iteration barrier overhead rather than
measuring the actual triad.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant